Back

Molecular Genetics and Genomics

Springer Science and Business Media LLC

All preprints, ranked by how well they match Molecular Genetics and Genomics's content profile, based on 12 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
WGBS of Differentiating Adipocytes Reveals Variations in DMRs and Context-Dependent Gene Expression

Yadav, B.; Singh, D.; Mantri, S.; Rishi, V.

2024-03-14 genomics 10.1101/2024.03.14.583264 medRxiv
Top 0.1%
6.7%
Show abstract

Obesity, characterised by the accumulation of excess fat, is a complex condition resulting from the combination of genetic and epigenetic factors. Recent studies have found correspondence between DNA methylation and cell differentiation, suggesting a role of the former in cell fate determination. There is a lack of comprehensive understanding concerning the underpinnings of preadipocyte differentiation, specifically when cells are undergoing terminal differentiation (TD). To gain insight into dynamic genome-wide methylation, 3T3 L1 preadipocyte cells were differentiated by a hormone cocktail. The genomic DNA was isolated from undifferentiated cells and 4 hrs (4H), 2 days (2D) post-differentiated cells, and 15 days (15D) TD cells. We employed whole-genome bisulfite sequencing (WGBS) to ascertain global genomic DNA methylation alterations at single base resolution as preadipocyte cells differentiate. The genome-wide distribution of DNA methylation showed similar overall patterns in pre- and post- and terminally differentiated adipocytes, according to WGBS analysis. DNA methylation decreases at 4H after differentiation initiation, followed by methylation gain as cells approach TD. Studies revealed novel differentially methylated regions (DMRs) associated with adipogenesis. DMR analysis suggested that though DNA methylation is global, noticeable changes are observed at specific sites known as hotspots. Hotspots are genomic regions rich in transcription factor (TF) binding sites and exhibit methylation-dependent TF binding. Subsequent analysis indicated hotspots as part of DMRs. The gene expression profile of key adipogenic genes in differentiating adipocytes is context-dependent, as we found a direct and inverse relationship between promoter DNA methylation and gene expression.

2
Genomic Insights of Bruneian Malays

Azmi, M.; Chen, L.; Idris, A.; Lu, Z. H.

2022-06-03 genomics 10.1101/2022.06.01.492266 medRxiv
Top 0.1%
6.1%
Show abstract

The Malays and their many sub-ethnic groups collectively make up one of the largest population groups in Southeast Asia. However, their genomes, especially those from Brunei, remain very much underrepresented and understudied. Here, we analysed the publicly available WGS and genotyping data of two and 39 Bruneian Malay individuals, respectively. NGS reads from the two individuals were first mapped against the GRCh38 human reference genome and their variants called. Of the total [~]5.28 million short nucleotide variants and indels identified, [~]217K of them were found to be novel; with some predicted to be deleterious and associated with risk factors of common non-communicable diseases in Brunei. Unmapped reads were next mapped against the recently reported novel Chinese and Japanese genomic contigs and de novo assembled. [~]227 Kbp genomic sequences missing in GRCh38 and a partial open reading frame encoding a potential novel small zinc finger protein were successfully discovered. Interestingly, although the Malays in Brunei, Singapore and Malaysia share >83% common variants, principal component and admixture analysis comparing the genetic structure of the local Malays against other Asian population groups suggested that they are genetically closer to some Filipino ethnic groups than the Malays in Malaysia and Singapore. Taken together, our work provides the first comprehensive insight into the genomes of the Bruneian Malay population.

3
Reassessing the association of VDR and its polymorphisms with tuberculosis in global populations.

Das, D.; Chaubey, G.

2023-12-11 genomics 10.1101/2023.12.09.570914 medRxiv
Top 0.1%
4.3%
Show abstract

BackgroundVitamin D is a hormone that regulates the calcium homeostasis of the body. Besides this classical function, it is also regarded as an important immunomodulator. Most active Vitamin D actions are mediated through the Vitamin D receptor (VDR), a transcription factor and also a member of the nuclear receptor superfamily. In this study, we explored the phylogeographic attributes of the four most well-known polymorphisms of the VDR gene namely rs7975232 (ApaI), rs731235 (TaqI), rs1544410 (BsmI), rs2228570 (FokI) and also evaluated their association with the incidence of tuberculosis in global populations. This study integrated several in-silico approaches on population databases to evaluate the pattern of distribution, linkage and selection patterns of these SNPs. ResultsThe ancestral alleles of rs7975232, rs731235, and rs1544410 are still present in over 50% frequency in modern human populations. These SNPs also have a very strong linkage disequilibrium among themselves in all population groups but no haplotype blocks are seen in South Asian populations constituting these polymorphisms. The selection results reveal a negative Tajimas D value in West and East Eurasian populations suggesting positive selection in these regions... In correlation studies, we found no association between the incidence of tuberculosis and the allele or genotype frequency of these four SNPs. ConclusionThe four SNPs of VDR behave differently in South Asian populations as compared to West and East Eurasian populations but no significant association was found with the incidence of tuberculosis in global populations.

4
Relationship among evolutionary distance, variance-covariance matrix and principal component analysis

Misawa, K.

2022-03-04 genomics 10.1101/2022.03.02.482744 medRxiv
Top 0.1%
4.3%
Show abstract

Principal component analyses (PCAs) are often used to visualize patterns of genetic variation in human populations. Previous studies showed a close correspondence between genetic and geographic distances. In such PCAs, the principal components are eigenvectors of the datas variance-covariance matrix, which is obtained by a genetic relationship matrix (GRM). However, it is difficult to apply GRM to multiallelic sites. In this paper, I showed that a PCA from GRM is equivalent to multidimensional scaling (MDS) from nucleotide differences. Therefore, a PCA can be conducted using nucleotide differences. The new method provided in this study provides a straightforward method to predict the effects of different demographic processes on genetic diversity.

5
The linear correlation between genome size and the size of the non-transcribing region

Chen, Z.-R.

2024-09-22 genomics 10.1101/2024.09.19.613789 medRxiv
Top 0.1%
4.1%
Show abstract

BackgroundThe genome sizes of organisms vary widely (C-value paradox). There are non-transcribing regions in the genome that neither encode proteins nor RNA entities. There are several hypotheses about the function of these regions: one suggests that they are unannotated functional areas, while another views them as genomic isolation zones that reduce mutations in coding regions. MethodStatistical analysis was conducted on the transcribing regions (including areas annotated as genes and transcribed pseudogenes) and non-transcribing regions, protein-coding regions (Coding sequence, CDS), and genome sizes using annotation files from 63,866 species genomes in the NCBI RefSeq database. ResultsThere is a significant linear relationship between the size of non-transcribing genomic regions and overall genome size across species, with varying proportional coefficients among different phyla (realms for viruses). As genome size increases, the proportion of non-transcribing regions gradually rises, eventually approaching a linear proportional limit, resembling one arm of hyperbolic functions. Eukaryotes show high linear correlation, with the highest in Streptophyta and the lowest in Apicomplexa. In eukaryotes, the size of the coding region increases with genome size, but the increasing trend diminishes (proportionally decreases). In non-eukaryotes, the size of the coding region maintains a linear relationship with genome size. ConclusionThe size of non-transcribing region in species may be subject to some strict quantitative control mechanism, showing that genome and non-transcribing genome sizes increase proportionally with the expansion of the transcribing genome, indicating a strict balance between expansion and energy conservation. The proportion of non-transcribed genomes in eukaryotes is conservative (although the sequences are not), and the presence of non-transcribing genomes has significant implications for the evolution or survival of species. Thus, I propose a new hypothesis about the non-transcribing genome, that it is a space for generating new genes from scratch, and the different proportional coefficients among phyla are due to their different positions in energy transfer. Graphic Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=104 SRC="FIGDIR/small/613789v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@dc3e88org.highwire.dtl.DTLVardef@18d70e8org.highwire.dtl.DTLVardef@efb92corg.highwire.dtl.DTLVardef@66068b_HPS_FORMAT_FIGEXP M_FIG C_FIG

6
Whole genome analysis of four Bangladeshi individuals

Khan, S.; Akter, S.; Goswami, B.; Habib, A.; Banu, T. A.; Barton, C.; Osman, E.; Samir, S.; Arjuman, F.; Hasan, S.; Hossain, M. M.

2020-05-23 genomics 10.1101/2020.05.21.109058 medRxiv
Top 0.1%
4.0%
Show abstract

Whole-genome sequencing (WGS) is a comprehensive method for analysing entire genomes and this has been instrumental in characterizing the single nucleotide polymorphisms associated with different diseases including cancer, diabetes, cardiovascular diseases and many others. In this paper we undertake a pilot study for sequencing four Bangladeshi individuals and profiling their single nucleotide variants. Our findings shed possible light on specific biological pathways effected by such variants in this population.

7
Correlation analysis among single nucleotide polymorphisms in thirteen language genes and culture/education parameters from twenty-six countries

Sun, B.; Guo, C.; Zhang, Z.

2021-08-24 evolutionary biology 10.1101/2021.08.22.457292 medRxiv
Top 0.1%
4.0%
Show abstract

Language is a vital feature of any human culture, but whether language gene polymorphisms have meaningful correlations with some cultural characteristics during the long-run evolution of human languages largely remains obscure (uninvestigated). This study would be an endeavor example to find evidences for the above questions answer. In this study, the collected basic data include 13 language genes and their randomly selected 111 single nucleotide polymorphisms (SNPs), SNP profiles, 29 culture/education parameters, and estimated cultural context values for 26 representative countries. In order to undertake principal component analysis (PCA) for correlation search, SNP genotypes, cultural context and all other culture/education parameters have to be quantitatively represented into numerical values. Based on the above conditions, this study obtained its preliminary results, the main points of which contain: (1) The 111 SNPs contain several clusters of correlational groups with positive and negative correlations with each other; (2) Low cultural context level significantly influences the correlational patterns among 111 SNPs in the principal component analysis diagram; and (3) Among 29 culture/education parameters, several basic characteristics of a language (the numbers of alphabet, vowel, consonant and dialect) demonstrate least correlations with 111 SNPs of 13 language genes.

8
Gene size matters: What determines gene length in the human genome?

Lopes, I.; Altab, G.; Raina, P.; de Magalhaes, J. P.

2020-01-10 genomics 10.1101/2020.01.10.901272 medRxiv
Top 0.1%
4.0%
Show abstract

While it is expected for gene length to be influenced by factors such as intron number and evolutionary conservation, we have yet to fully understand the connection between gene length and function in the human genome. In this study, we show that, as expected, there is a strong positive correlation between gene length and the number of SNPs, introns and protein size. Amongst tissue specific genes, we find that the longest genes are expressed in blood vessels, nerve, thyroid, cervix uteri and brain, while the smallest genes are expressed within the pancreas, skin, stomach, vagina and testis. We report, as shown previously, that natural selection suppresses changes for genes with longer lengths and promotes changes for smaller genes. We also observed that longer genes have a significantly higher number of co-expressed genes and protein-protein interactions. In the functional analysis, we show that bigger genes are often associated with neuronal development, while smaller genes tend to play roles in skin development and in the immune system. Furthermore, pathways related to cancer, neurons and heart diseases tend to have longer genes, with smaller genes being present in pathways related to immune response and neurodegenerative diseases. We hypothesise that longer genes tend to be associated with functions that are important early in life, while smaller genes play a role in functions that are important throughout the organisms whole life, like the immune system which require fast responses. Author SummaryEven though the human genome has been fully sequenced, we still do not fully grasp all of its nuances. One such nuance is the length of the genes themselves. Why are certain genes longer than others? Is there a common function shared by longer/smaller genes? What exactly makes gene longer? We tried answering these questions using a variety of analysis. We found that, while there was not a particular strong factor in genes that influenced their size, there could be an influence of several gene characteristics in determining the length of a gene. We also found that longer genes are linked with the development of neurons, cancer, heart diseases and muscle cells, while smaller genes seem to be mostly related with the immune system and the development of the skin. This led us to believe that, whether the gene has an important function early in our life, or throughout our whole lives, or even if the function requires a rapid response, that its gene size will be influenced accordingly.

9
Genetic variants of human platelet antigens in the Indian population from 1029 whole genomes

Rophina, M.; Bhoyar, R. C.; Imran, M.; Senthivel, V.; Divakar, M. K.; Mishra, A.; Jolly, B.; Sivasubbu, S.; Scaria, V.

2022-10-31 genomics 10.1101/2022.10.29.514338 medRxiv
Top 0.1%
3.6%
Show abstract

BackgroundGenetic variants in human platelet antigens (HPAs) considered as allo- or auto antigens are associated with various disorders including neonatal alloimmune thrombocytopenia, platelet transfusion refractoriness and post-transfusion purpura. While global differences in genotype frequencies were observed, the distribution of HPA variants in the Indian population are largely unknown. This study aims to explore the landscape of HPA variants in India to provide a basis for risk assessment and management of related complications. Materials and methodsPopulation specific frequencies of genetic variants associated with the 35 classes of HPAs (HPA-1 to HPA-35) were estimated by systematically analyzing genomic variations of 1029 healthy Indian individuals as well as from global population genome datasets.. ResultsAllele frequencies of the most clinically relevant HPA systems in the Indian population were found as follows, HPA-1a - 0.89, HPA-1b - 0.15, HPA-2a - 0.94, HPA-2b - 0.05, HPA-3a - 0.66, HPA-3b - 0.36, HPA-4a - 1.00, HPA-4b - 0, HPA-5a - 0.92, HPA-5b - 0.08, HPA-6a - 1.00, HPA-6b - 0, HPA-15a - 0.58 and HPA-15b - 0.42. In addition, HPA-4b allele frequencies were found to be significantly higher in India in comparison to global populations. ConclusionThis study provides the first comprehensive analysis of HPA allele and genotype frequencies using large scale representative whole genome sequencing data of the Indian population.

10
Gene Expression and Physiological traits in Mice

Cruz, I. N.; Ramos, R. B.

2023-06-01 genomics 10.1101/2023.05.30.542939 medRxiv
Top 0.1%
3.3%
Show abstract

BackgroundGene expression regulates several complex traits observed. In this study, datasets comprising of transcriptome information and clinical traits regarding fat composition and vitals were analyzed via several statistical methods in order to find relations between genes and clinical outcomes. ResultsBiological big data is diverse and numerous, which makes for a complex case study and difficulties to stablish a metric. Histological data with semi-quantitative scores proved unreliable to correlate with other vitals, such as cholesterol composition, which complicates prediction of clinical outcomes. A composition of vitals, turned out to be a better variable for regression and factors for gene analysis. Several genes were found to be statistically significant after statistical analysis by ANOVA regarding the progressive categories of the preferred clinical variable. ConclusionsANOVA is proposed as a method for genetic information retrieval in order to extract biological meaning from RNA seq or microarray data, accounting for multiple classes of target variables. It Provides a reliable statistical method to associate genes or clusters of genes with particular traits. Supplementary informationSupplementary data are available in annexes.

11
Population genomics of a natural Cannabis sativa L. collection from Iran identifies novel genetic loci for flowering time, morphology, sex and chemotyping

Dehnavi, M. M.; Damerum, A.; Taheri, S.; Ebadi, A.; Panahi, S.; Hodgin, G.; Brandley, B.; Salami, S. A.; Taylor, G.

2024-05-10 genomics 10.1101/2024.05.07.593022 medRxiv
Top 0.1%
3.3%
Show abstract

Future breeding and selection of Cannabis sativa L. for drug production and industrial purposes require a source of germplasm with wide genetic variation, such as that found in wild relatives and progenitors of highly cultivated plants. Limited directional selection and breeding have occurred in this crop, especially informed by molecular markers. Here, we investigated the population genomics of a natural cannabis collection of male and female individuals from differing climatic zones in Iran. Using Genotyping-By-Sequencing (GBS), we sequenced 228 genotypes from 35 populations. The results obtained from GBS were used to perform association analysis identifying links between genotype and important phenotypes, including inflorescence characteristics, flowering time, plant morphology, tetrahydrocannabinol (THC) content, cannabidiol (CBD) content and sex. Approximately 23,266 significant SNPs of high quality were detected to establish associations between markers and traits, and population structure showed that Iranian cannabis plants fall into five groups. A comparison of Iranian samples from this study to global data suggests that the Iranian population is distinctive and, in general, is closer to marijuana than to hemp, although some populations in this collection are closer to hemp. The GWAS results showed that novel genetic loci, not previously identified, contribute to sex, yield and chemotype traits in cannabis and are worthy of further study.

12
Identification of Essential Temperature-Stressed Genes From Apis mellifera Hypopharyngeal Glands Transcriptomes Under Variable Temperatures

Maigoro, A. Y.; Lee, J. H.; Lee, S.; Kwon, H. W.

2023-11-22 genomics 10.1101/2023.11.22.568201 medRxiv
Top 0.1%
3.2%
Show abstract

Temperature is one of the essential abiotic factors required for honey bee survival and pollination. It affects honey bee physiology, behavior, and expression of related genes. Also, considered one of the major factors contributing to colony collapse disease (CCD). In this research, RNA-seq analysis was performed using hypopharyngeal glands (HGs) tissue at low (18 {degrees}C), high (25 {degrees}C), and regular (22 {degrees}C) temperatures. Differentially expressed genes (DEGs) were identified after comparing the three groups with one another based on temperature differences using DESeq analysis. 5196 common DEGs (cDEGs) were identified among the groups. They are highly enriched in RNA processing and RNA metabolism process while the KEGG pathway enrichment analysis showed that the cDEGs are enriched in longevity regulating pathway, MAPK signaling pathway-fly, and Glycerophospholipid metabolism. Further, 360 temperature-stressed genes identified are highly enriched in translation, oxidative activity, and ribonucleoprotein complex. The enriched KEGG pathway includes ribosome, oxidative phosphorylation, fatty acid metabolism, and citrate cycle (TCA cycle). All the top ten (10) hub genes among the 360 temperature-stressed genes are found up-regulated. In addition, heat-shock protein 90 (HSP90) known as the stressed response gene, and Gr10, the amino acid response gene were up-regulated and down-regulated respectively in the temperature-stressed group. Low expression of Gr10 under temperature-stress can affect nursing behavior and bee development. Ultimately, these findings will help in identifying honeybee-temperature survival mechanisms under varying temperature effects.

13
Bifidobacterium is enriched in gut microbiome of Kashmiri women with polycystic ovary syndrome

Hassan, S.; Kaakinen, M. A.; Draisma, H.; Ganie, M. A.; Balkhiyarova, Z.; Vogazianos, P.; Shammas, C.; Selvin, J.; Antoniades, A.; Demirkan, A.; Prokopenko, I.

2019-07-30 genomics 10.1101/718510 medRxiv
Top 0.1%
3.2%
Show abstract

Polycystic ovary syndrome (PCOS) is a common endocrine condition in women of reproductive age understudied in non-European populations. In India, PCOS affects the life of up to 19.4 million women of age 14-25 years. Gut microbiome composition might contribute to PCOS susceptibility. We profiled the microbiome in DNA isolated from faecal samples by 16S rRNA sequencing in 19/20 women with/without PCOS from Kashmir, India. We assigned genera to sequenced species with an average 121k reads depth and included bacteria detected in at least 1/3 of the subjects or with average relative abundance [≥]0.1%. We compared the relative abundances of 40/58 operational taxonomic units in family/genus level between cases and controls, and in relation to 33 hormonal and metabolic factors, by multivariate analyses adjusted for confounders, and corrected for multiple testing. Seven genera were significantly enriched in PCOS cases: Sarcina, Alkalibacterium and Megasphaera, and previously reported for PCOS Bifidobacterium, Collinsella, Paraprevotella and Lactobacillus. We identified significantly increased relative abundance of Bifidobacteriaceae (median 6.07% vs. 2.77%) and Aerococcaceae (0.03% vs. 0.004%), whereas we detected lower relative abundance Peptococcaceae (0.16% vs. 0.25%) in PCOS cases. For the first time, we identified a significant direct association between butyrate producing Eubacterium and follicle-stimulating hormone levels. We observed increased relative abundance of Collinsella and Paraprevotella with higher fasting blood glucose levels, and Paraprevotella and Alkalibacterium with larger hip and waist circumference, and weight. We show a relationship between gut microbiome composition and PCOS linking it to specific reproductive health metabolic and hormonal predictors in Indian women.

14
Red cell enzyme polymorphisms in Muslim population of Eastern UP, India

Kumar, P.; Rai, V.

2022-11-28 genomics 10.1101/2022.11.24.517832 medRxiv
Top 0.1%
3.2%
Show abstract

BackgroundRed cell enzyme polymorphisms have been used to study genetic variations in several human populations/countries worldwide. Owing to considerable ethnic and cultural heterogeneity in the Indian population, it is imperative to study different caste and religion of different regions. Aims and ObjectiveTo determine allele frequency of five Red Cell Enzymes (ADA, AK1, ESD, GLO1) in Muslim population of Eastern Uttar Pradesh Materials and methodsBlood samples were collected from 200 unrelated individuals belonging to Muslim community of eastern Uttar Pradesh. The phenotypes of ADA, AK1, ESD and GLO1 systems were determined by agarose/starch gel electrophoresis. The allele frequencies were calculated by gene count method. ResultsThe calculated frequencies of the alleles are as follows: ADA*1= 0.907, ADA*2= 0.092; AK*1= 0.92, AK*2= 0.08; ESD*1= 0.765, ESD*2= 0.235; GLO*1= 0.27,GLO*2 =0.73 ConclusionThe comparison of the allele frequencies of the four RBC enzymes studied in the present report with those of Asian populations showed that the allele frequencies are close to other Asian populations.

15
Statistical analysis of number of genes and chromosome lengths of different microbial species

Ng, W.

2022-08-16 genomics 10.1101/2022.08.13.503871 medRxiv
Top 0.1%
3.2%
Show abstract

Genome architecture concerns the organisation of genes on a chromosome, and has important implications to the fidelity in which genes are encoded on the chromosome, and how the information is read by DNA polymerase and RNA polymerase. This facet of genomics did receive attention in the early epoch of genomics, but it has received less attention in contemporary genomics as attention shifts to structural and functional genomics with the goal of annotating the function of each gene in the genome. This work sought to uncover relationships between number of genes and chromosome length in a variety of bacteria and archaea species as a preamble to understanding the prevalence and importance of repetitive sequences in the genome of prokaryotic species. Aggregate results with the ensemble of prokaryotic species profiled revealed a positive linear correlation between number of genes and chromosome length. Upon dissection into the Bacteria and Archaea domains, the linear relationship described above still stands for Bacteria but starts to break down in Archaea. This suggests that repetitive sequences are more important to Archaea species, which generally have a smaller genome (1.8 to 2.8 Mbp) and fewer genes (1500 to 2400) compared to bacterial species. In comparison, the bacterial genome is larger (4 to 5.6 Mbp), and encodes more genes (3600 to 5100). Overall, the results highlight that bacterial genome are efficiently encoded with few repetitive sequences. This, however, is not true for archaeal genome, which provides another line of evidence supporting the notion that archaea are ancestral eukaryotic cells, which like the archaea also houses large repetitive sequences. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=141 SRC="FIGDIR/small/503871v1_ufig1.gif" ALT="Figure 1"> View larger version (15K): org.highwire.dtl.DTLVardef@1632506org.highwire.dtl.DTLVardef@13e91forg.highwire.dtl.DTLVardef@12e1316org.highwire.dtl.DTLVardef@1e7381e_HPS_FORMAT_FIGEXP M_FIG C_FIG Short descriptionStatistical analysis across an ensemble of 59 microbial species revealed a strong linear correlation between number of genes and chromosome length. This suggests that prokaryotic genomes are highly compact with genes, and do not carry significant amounts of repeats unlike the case in eukaryotic organisms. The result holds significant implications for our understanding of genome evolution and compaction in prokaryotic organisms, and what drove their accession as foundational species of many ecosystems. Subject areasgenomics, molecular biology, evolutionary biology, bioinformatics, systems biology,

16
Phyloepigenetics in phylogeny analyses

Santourlidis, S.

2024-08-18 evolutionary biology 10.1101/2024.08.14.607911 medRxiv
Top 0.1%
3.2%
Show abstract

Long-standing, continuous blurring and controversies in the field of phylogenetic interspecies relations, associated with insufficient explanations for dynamics and variability of speeds of evolution in mammals, hint to a crucial missing link. It has been suggested that transgenerational epigenetic inheritance and the concealed mechanisms behind play a distinct role in mammalian evolution. Here, a comprehensive sequence alignment approach in hominid species, i.e., Homo sapiens, Homo neanderthalensis, denisovan human, Pan troglodytes, Pan paniscus, Gorilla gorilla and Pongo pygmaeus, comprising conserved CpG islands of housekeeping genes, uncover evidence for a distinct variability of CpG dinucleotides. Applying solely these evolutionary consistent and inconsistent CpG sites in a classic phylogenetic analysis, calibrated by the divergence time point of the common chimpanzee (Pan troglodytes) and the bonobo or pygmy chimpanzee (Pan paniscus), a "phylo-epigenetic" tree has been generated which precisely recapitulates branch points and branch lengths, i.e., divergence events and relations, as they have been broadly suggested in the current literature, based on comprehensive molecular phylogenomics and fossil records. I suggest here that CpG dinucleotides changes at CpG islands are of superior importance for evolutionary development and determine the emerging DNA methylation profiles.

17
Invariant Genes in Human Genomes

Pathak, A. K.; Jainarayanan, A. K.; Brahmachari, S. K.

2019-08-20 genomics 10.1101/739706 medRxiv
Top 0.1%
3.2%
Show abstract

With large-scale human genome and exome sequencing, a lot of focus has gone in studying variations present in genomes and their associations to various diseases. Since major emphasis has been put on their variations, less focus has been given to invariant genes in the population. Here we present 60,706 genomes from the ExAC database to identify population specific invariant genes. Out of 1,336 total genes drawn from various population specific invariant genes, 423 were identified to be mostly (allele frequency less than 0.001) invariant across different populations. 46 of these invariant genes showed absolute invariance in all populations. Most of these common invariant genes have homologs in primates, rodents and placental mammals while 8 of them were unique to human genome and 3 genes still had unknown functions. Surprisingly, a majority were found to be X-linked and around 50% of these genes were not expressed in any tissues. The functional analysis showed that the invariant genes are not only involved in fundamental functions like transcription and translation but also in various developmental processes. The variations in many of these invariant genes were found to be associated with cancer, developmental diseases and dominant genetic disorders.

18
Identification of potential biomarkers associated with pathogenesis of primary prostate cancer based on meta-analysis approaches

Sepahi, N.; Piran, M.; Piran, M.; Ghanbariasad, A.

2020-03-05 genomics 10.1101/2020.03.05.978205 medRxiv
Top 0.1%
3.2%
Show abstract

Worldwide prostate cancer (PCa) is recognized as the second most common diagnosed cancer and the fifth leading cause of cancer death among men globally. Rising incidence rates of PCa have been observed over the last few decades. It is necessary to improve prostate cancer detection, diagnosis, treatment and survival. However, there are few reliable biomarkers for early prostate cancer diagnosis and prognosis. In the current study, systems biology method was applied for transcriptomic data analysis to identify potential biomarkers for primary PCa. We firstly identified differentially expressed genes (DEGs) between primary PCa and normal samples. Then the DEGs were mapped in Wikipathways and gene ontology database to conduct functional categories enrichment analysis. 1575 unique DEGs with adjusted p-value < 0.05 were achieved from two sets of DEGs. 132 common DEGs between two sets of DEGs were retrieved. The final DEGs were selected from 60 common upregulated and 72 common downregulated genes between datasets. In conclusion, we demonstrated some potential biomarkers (FOXA1, AGR2, EPCAM, CLDN3, ERBB3, GDF15, FHL1, NPY, DPP4, and GADD45A) and HIST2H2BE as a candidate one which are tightly correlated with the pathogenesis of PCa.

19
Long antiparallel open reading frames are unlikely to be encoding essential proteins in prokaryotic genomes

Moshensky, D.; Alexeevski, A.

2019-08-05 genomics 10.1101/724807 medRxiv
Top 0.1%
3.1%
Show abstract

The origin and evolution of genes that have common base pairs (overlapping genes) are of particular interest due to their influencing each other. Especially intriguing are gene pairs with long overlaps. In prokaryotes, co-directional overlaps longer than 60 bp were shown to be nonexistent except for some instances. A few antiparallel prokaryotic genes with long overlaps were described in the literature. We have analyzed putative long antiparallel overlapping genes to determine whether open reading frames (ORFs) located opposite to genes (antiparallel ORFs) can be protein-coding genes.\n\nWe have confirmed that long antiparallel ORFs (AORFs) are observed reliably to be more frequent than expected. There are 10 472 000 AORFs in 929 analyzed genomes with overlap length more than 180 bp. Stop codons on the opposite to the coding strand are avoided in 2 898 cases with Benjamini-Hochberg threshold 0.01.\n\nUsing Ka/Ks ratio calculations, we have revealed that long AORFs do not affect the type of selection acting on genes in a vast majority of cases. This observation indicates that long AORFs translations commonly are not under negative selection.\n\nThe demonstrative example is 282 longer than 1 800 bp AORFs found opposite to extremely conserved dnaK genes. Translations of these AORFs were annotated \"glutamate dehydrogenases\" and were included into Pfam database as third protein family of glutamate dehydrogenases, PF10712. Ka/Ks analysis has demonstrated that if these translations correspond to proteins, they are not subjected by negative selection while dnaK genes are under strong stabilizing selection. Moreover, we have found other arguments against the hypothesis that these AORFs encode essential proteins, proteins indispensable for cellular machinery.\n\nHowever, some AORFs, in particular, dnaK related, have been found slightly resisting to synonymous changes in genes. It indicates the possibility of their translation. We speculate that translations of certain AORFs might have a functional role other than encoding essential proteins.\n\nEssential genes are unlikely to be encoded by AORFs in prokaryotic genomes. Nevertheless, some AORFs might have biological significance associated with their translations.\n\nAuthor summaryGenes that have common base pairs are called overlapping genes. We have examined the most intriguing case: if gene pairs encoded on opposite DNA strands exist in prokaryotes. An intersection length threshold 180 bp has been used. A few such pairs of genes were experimentally confirmed.\n\nWe have detected all long antiparallel ORFs in 929 prokaryotic genomes and have found that the number of open reading frames, located opposite to annotated genes, is much more than expected according to statistical model. We have developed a measure of stop codon avoidance on the opposite strand. The lengths of found antiparallel ORFs with stop codon avoidance are typical for prokaryotic genes.\n\nComparative genomics analysis shows that long antiparallel ORFs (AORFs) are unlikely to be essential protein-coding genes. We have analyzed distributions of features typical for essential proteins among formal translations of all long AORFs: prevalence of negative selection, non-uniformity of a conserved positions distribution in a multiple alignment of homologous proteins, the character of homologs distribution in phylogenetic tree of prokaryotes. All of them have not been observed for the majority of long AORFs. Particularly, the same results have been obtained for some experimentally confirmed AOGs.\n\nThus, pairs of antiparallel overlapping essential genes are unlikely to exist. On the other hand, some antiparallel ORFs affect the evolution of genes opposite that they are located. Consequently, translations of some antiparallel ORFs might have yet unknown biological significance.

20
Genome-wide characterization and identification of synonymous codon usage patterns in Plasmodium knowlesi

Yadav, M. K.; Gajbhiye, S.

2021-01-02 genomics 10.1101/2021.01.01.425038 medRxiv
Top 0.1%
2.9%
Show abstract

Codon usage bias is a ubiquitous phenomenon occurring at both, interspecies and intraspecies level in different organisms. P. knowlesi, whose natural host is long-tailed Macaque monkeys, has recently started infecting humans as well. The genome as well as coding sequence data of P. knowlesi is used to understand their codon usage pattern in the light of other human infecting Plasmodium species: P. vivax and P. falciparum. The different codon usage indicators: GC content, relative synonymous codon usage, effective number of codon and codon adaptation index are studied to analyze codon usage in the Plasmodium species. The codon usage pattern is found to be less conserved in studied Plasmodium species, and changes species to species at the genus level. The codon usage pattern of P. knowlesi shows similarity to P. vivax as compared to P. falciparum. The ENC vs. GC3 study indicates that compositional constraints and translation selection is the decisive forces responsible for shaping their codon usage. The studies Plasmodium species shows a higher usage of A/T ending optimal codons. This favors the codon bias in P. knowlesi and P. vivax is due to high selection pressure and in P. falciparum, the compositional mutational pressure is a dominant force. In a nutshell, our finding suggests that the more or less similar codon usage pattern of P. knowlesi and P. vivax may suggest the similar host invasion and immune evasion strategies for disease establishment.